In this lesson
Phase 5 · Lesson 5.1

AI Ethics: Bias, Fairness and Accountability

The most technically impressive AI system can still cause serious harm. Understanding where things go wrong and why is now a core responsibility for every AI practitioner.

🕑 35 min read ⚖ 4 real case studies ⚖ No code required

Why AI Ethics Matters Now

For much of its early history, AI was a research discipline with limited real-world reach. That has changed dramatically. AI systems now decide who receives a loan, who is flagged as a flight risk in court, whose job application advances, and which patients are prioritised for additional medical care. When these systems fail or discriminate, the consequences are not abstract. They are experienced by real people.

The field of AI ethics is not about making AI "nice." It is about ensuring that the power of these systems is exercised responsibly, that harms are identified before deployment rather than discovered after, and that the people most affected by AI decisions have meaningful recourse when things go wrong.

This is not a soft skill or an optional extra. It is increasingly a legal requirement, a professional expectation, and a practical necessity for building systems that people will actually trust and use.

A Common Misconception

Many people assume that because an AI system is trained on data rather than programmed with human prejudices, it must be objective. This is wrong. Data is generated by human societies, and human societies carry historical inequalities. A model trained on biased data learns those biases. Automation does not neutralise discrimination. In some cases, it amplifies it and makes it harder to challenge.

Where Bias Comes From

Bias in AI systems does not usually arise from malicious intent. It enters the pipeline in several distinct, often invisible ways. Understanding where it enters is the first step to addressing it.

The Bias Pipeline: Where Bias Enters AI Systems Data Collection Historical bias Underrepresentation Labelling & Annotation Annotator bias Measurement error Model Training Objective mismatch Proxy discrimination Evaluation & Testing Non-representative benchmarks Deployment & Feedback Context drift Feedback loops Biased outputs become future training data, amplifying the original bias

Bias can enter at every stage of the AI pipeline. Addressing it at deployment alone is too late. Rigorous checks are needed at each step, with particular attention to the feedback loop that can compound errors over time.

Historical Bias

The world has not always been fair. Training on historical records encodes historical inequalities. A hiring model trained on past hiring decisions inherits the biases of past hiring managers.

Representation Bias

When training data underrepresents certain groups (women in STEM datasets, darker skin tones in medical imaging), the model performs worse for those groups.

Measurement Bias

The variables used to measure a concept may not measure it equally across groups. Using healthcare cost as a proxy for healthcare need systematically underestimates the needs of groups who historically received less care.

Feedback Loop Bias

A predictive policing model directs more patrols to certain neighbourhoods, leading to more arrests there, which generates more data confirming the model's predictions. The bias becomes self-reinforcing.

Aggregation Bias

Training a single model on a diverse population can obscure important subgroup differences. A medical model trained on combined data may perform well on average but fail for specific demographic groups.

Deployment Bias

A model is used in a context it was not designed for. A risk score developed for one country's legal system is applied in another with very different social conditions and base rates.

Real Case Studies

Examining specific documented cases is essential. Abstract discussions of bias are less instructive than concrete examples of how it has caused real harm.

Case Study 1 · Criminal Justice

COMPAS Recidivism Risk Tool

COMPAS (Correctional Offender Management Profiling for Alternative Sanctions) is a commercial tool used by courts across the United States to predict the likelihood that a defendant will re-offend. In 2016, investigative journalists at ProPublica published an analysis of 7,000 defendants in Broward County, Florida.

Their analysis found that the tool was significantly more likely to incorrectly label Black defendants as high risk when they did not go on to re-offend, and more likely to incorrectly label white defendants as low risk when they did. The tool's overall accuracy was similar across racial groups, but the types of errors were distributed very differently.

Northpointe, the company behind COMPAS, responded that the tool satisfied a different statistical definition of fairness: that its risk scores meant the same thing regardless of race (that is, a score of 7 corresponded to the same re-offending rate for Black and white defendants). Both claims were true simultaneously. This case launched a now-famous debate in the research community about which mathematical definition of fairness is most appropriate, and whether they can all be satisfied at once when re-offending rates differ between groups.

The lesson: accuracy parity at the overall level can mask very different error distributions across groups. The choice of which fairness metric to optimise is not a technical question. It is a value judgment that should involve affected communities.
Case Study 2 · Hiring

Amazon's Automated Recruiting Tool

Amazon developed an AI recruiting tool that scored job applicants on a scale of one to five stars. The tool was trained on resumes submitted to Amazon over a ten-year period, a period during which the tech industry was overwhelmingly male. The model learned to penalise resumes containing the word "women's" (as in "women's chess club") and downgraded graduates of all-women's colleges.

Amazon's team discovered these issues in 2015, made corrections, but found they could not guarantee the tool would not find other problematic proxies for gender. The project was scrapped in 2018. Reuters reported on the case publicly that year.

The lesson: even when protected attributes like gender are excluded from training features, a model can learn to use proxies for those attributes. The absence of a sensitive variable in the feature set does not guarantee fairness.
Case Study 3 · Healthcare

Racial Bias in a Medical Care Algorithm

A 2019 study published in Science by Obermeyer et al. analysed a widely deployed commercial algorithm used by US health systems to identify patients who needed extra care management. The algorithm used predicted healthcare costs as a proxy for healthcare need. The researchers found that for the same healthcare cost level, Black patients were significantly sicker than white patients.

The root cause was a measurement bias: Black patients historically received less care for the same conditions, so their costs were lower even though their underlying needs were the same. The algorithm interpreted lower costs as lower need, and as a result, at any given risk score threshold, Black patients who were enrolled in care management programmes were considerably sicker than their white counterparts. The authors estimated this reduced the fraction of Black patients receiving additional care by more than half compared to what an unbiased algorithm would have produced.

The lesson: choosing a proxy variable that seems reasonable (cost as a measure of need) can encode systemic inequalities in a way that is invisible until you specifically look for disparate impact across groups.
Case Study 4 · Computer Vision

Gender Shades: Disparities in Face Analysis

Joy Buolamwini (MIT Media Lab) and Timnit Gebru (Microsoft Research at the time) published "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification" at the 2018 ACM Conference on Fairness, Accountability, and Transparency. The study evaluated three major commercial gender classification systems from IBM, Microsoft, and Face++.

Across all three systems, overall accuracy ranged from 93.7% to 99.1%. But when broken down by skin tone and gender together, darker-skinned female subjects were classified correctly only 65.3% to 79.9% of the time. The disparity was intersectional: darker-skinned women fared far worse than lighter-skinned men or either group considered alone. The training datasets used by these systems contained disproportionately more lighter-skinned subjects.

After the study's publication, all three companies updated their systems and reported substantially improved performance across demographic groups. The study directly motivated Amazon, IBM, and Microsoft to later pause or limit the sale of facial recognition technology to police departments while accuracy and fairness standards were being debated.

The lesson: aggregate performance metrics hide subgroup disparities. Intersectional analysis (examining the interaction between multiple identity dimensions) reveals harms that group-by-group analysis can miss.

Fairness is Not One Thing

The COMPAS case highlighted a genuine mathematical tension. Researchers Chouldechova (2017) and Kleinberg, Mullainathan, and Raghavan (2016) independently proved that several commonly used fairness criteria are mathematically incompatible with each other whenever the positive outcome rate differs between groups. This is sometimes called the impossibility theorem of fairness.

There is no single correct definition of "fair" that a model can simultaneously satisfy. Different definitions reflect different values, and the choice between them has significant ethical implications.

Fairness Definition What It Requires When It Prioritises
Demographic Parity
(Statistical Parity)
Equal positive prediction rates across all groups. If 30% of Group A receives a loan, 30% of Group B should too. Equal representation of outcomes across groups, regardless of underlying differences in the predictor variables.
Equalized Odds Equal true positive rates AND equal false positive rates across groups. The model should be equally accurate for everyone. Equal quality of predictions. If you are incorrectly flagged as high risk, that probability should not depend on your group membership.
Calibration
(Predictive Parity)
A predicted probability of 70% should correspond to a 70% actual outcome rate for all groups. Scores mean the same thing across groups. Consistency of interpretation across groups. This is what Northpointe argued COMPAS satisfied.
Individual Fairness Similar individuals should receive similar predictions, regardless of group membership. Treating each person based on their own features rather than their demographic group average.
Counterfactual Fairness Would the prediction change if the individual's protected attribute changed, all else being equal? Causal reasoning about how the model uses protected attributes or their proxies.
The Core Tension

When the base rates of an outcome differ between groups (for example, if recidivism rates differ between demographic groups for historical reasons), you cannot simultaneously satisfy calibration and equalized false positive rates. You must choose. This is not a technical problem that better algorithms will solve. It is a societal question about which type of error is more acceptable, and who bears the cost of being wrong. The answer should come from democratic deliberation and input from affected communities, not from engineers optimising a loss function.

Explainability and Accountability

When an AI system makes a decision that affects someone's life, that person has a legitimate interest in understanding why. This is not just a philosophical point. The EU's General Data Protection Regulation (GDPR) includes provisions granting individuals rights related to automated decision-making, requiring that decisions be explainable in meaningful terms.

Interpretable Models vs Post-Hoc Explanations

Some models are inherently interpretable. A logistic regression or a shallow decision tree can be read directly by a human expert. The model's reasoning is transparent by design. Cynthia Rudin (Duke University) has argued forcefully, particularly in a 2019 paper in Nature Machine Intelligence, that in high-stakes domains (criminal justice, medical diagnosis, credit lending), we should prefer inherently interpretable models over complex black boxes with post-hoc explanations, because post-hoc explanations are approximations that may not accurately represent what the model actually computed.

Deep neural networks are not inherently interpretable. Their reasoning is distributed across millions of parameters. To explain them, researchers have developed post-hoc explanation methods:

SHAP (2017)

SHapley Additive exPlanations, developed by Lundberg and Lee. Uses game theory (Shapley values from cooperative game theory) to assign each input feature a contribution score for a specific prediction. Theoretically grounded and consistent.

LIME (2016)

Local Interpretable Model-agnostic Explanations, by Ribeiro, Singh, and Guestrin. Approximates the complex model locally around a specific prediction using a simple linear model. Intuitive but the local approximation may not always be faithful.

Grad-CAM (2017)

Gradient-weighted Class Activation Mapping, by Selvaraju et al. For image classification CNNs, produces a heatmap showing which regions of the image most influenced the prediction. Widely used in medical imaging to verify model reasoning.

Attention Visualisation

For Transformer models, plotting attention weights shows which tokens a model focused on when generating an output. Useful but research has shown attention weights do not always correspond to what actually drives predictions.

An Important Caveat on Explanations

Post-hoc explanations are approximations, not ground truth. A SHAP explanation tells you how a simplified model accounts for the prediction, not necessarily how the actual model computed it. Using these tools uncritically, or communicating them to users as if they were definitive, can create a false sense of transparency. Explanations are a starting point for scrutiny, not a substitute for it.

Regulation and Governance

Governments and international bodies have begun to respond to the risks of AI with regulatory frameworks. Understanding the landscape is becoming an important part of AI practice, particularly for anyone building systems that will be deployed commercially.

EU AI Act (2024)

The world's first comprehensive AI law, entering into force August 2024. It takes a risk-based approach: systems are classified by risk level (unacceptable, high, limited, minimal). High-risk systems (credit scoring, recruitment, critical infrastructure) face strict obligations including conformity assessments, transparency requirements, and human oversight mandates. AI systems that pose unacceptable risks (social scoring by governments, real-time remote biometric surveillance in public) are banned.

EU GDPR (2018)

While not AI-specific, the General Data Protection Regulation contains provisions directly relevant to AI: Article 22 gives individuals the right not to be subject to solely automated decisions with significant effects, and the right to a meaningful explanation. It also requires data minimisation, purpose limitation, and privacy by design. GDPR applies to any organisation processing data of EU residents.

UNESCO Recommendation (2021)

The first global standard on AI ethics, adopted by all 193 UNESCO member states. It covers human rights, transparency, accountability, fairness, environmental sustainability, and gender equality in AI. It is non-binding but represents the broadest international consensus on AI ethics principles.

US Executive Order on AI (2023)

Signed October 2023, it directed federal agencies to establish standards for AI safety and security, protect privacy, promote equity and civil rights, and support workers affected by AI. It required developers of powerful AI systems to share safety test results with the government. The Biden-era order was partially rescinded in early 2025, reflecting ongoing policy debates about AI governance in the US.

Beyond government regulation, voluntary frameworks have emerged from the research community. Model Cards (proposed by Timnit Gebru, Margaret Mitchell et al. in 2018) are structured documentation forms for ML models that disclose performance characteristics across demographic subgroups. Datasheets for Datasets (Gebru et al., 2021) provide an analogous framework for documenting training data. These tools have been widely adopted by major AI labs as part of responsible release practices.

What You Can Do as an AI Practitioner

Ethics is not only for policymakers and researchers. Every person who builds, deploys, or recommends AI systems has a responsibility to consider these issues. Here are concrete steps you can take at each stage of an AI project.

01

Audit your training data

Before training, examine the demographic composition of your dataset. Ask: who is represented? Who is not? What are the base rates of the outcome variable across subgroups? Document your findings.

02

Disaggregate your evaluation metrics

Never report only overall accuracy. Compute precision, recall, and F1 separately for each demographic subgroup relevant to your deployment context. Disparities that aggregate metrics hide become visible this way.

03

Question your proxy variables

Ask whether the variable you are using to measure the concept you care about actually measures it equally across groups. If it does not, consider whether an alternative measurement is available.

04

Involve affected communities

The people most likely to be impacted by your system should have input into its design, evaluation, and deployment conditions. This is not just ethical. It is also practically useful for identifying failure modes.

05

Document your models

Use a structured format like a Model Card to record what your model was trained on, what it was evaluated on, its intended use cases, its known limitations, and its performance across demographic groups.

06

Design for human oversight

For high-stakes decisions, build in meaningful human review rather than full automation. Ensure that humans in the loop have sufficient information, time, and authority to actually override the system when needed.

Further Reading

📚
Weapons of Math Destruction — Cathy O'Neil (2016)
Accessible, well-documented survey of how algorithms cause harm in education, finance, policing, and more. Essential reading for any AI practitioner.
📚
Algorithms of Oppression — Safiya Umoja Noble (2018)
Examines how search engine results can reflect and reinforce racial and gender bias, with broader lessons for AI fairness.
📚
Atlas of AI — Kate Crawford (2021)
Examines the material, political, and environmental costs of AI systems. Broadens the ethical lens beyond algorithmic fairness to include labour, power, and extraction.
📄
"Gender Shades" — Buolamwini and Gebru (FAccT 2018)
The original paper documenting intersectional disparities in commercial face analysis systems. Freely available online.
📄
"Dissecting Racial Bias in an Algorithm Used to Manage Health" — Obermeyer et al. (Science, 2019)
The healthcare cost proxy study. A model paper for how to audit an algorithm for disparate impact.

Key Takeaways

Coming Up: Large Language Models

In Lesson 5.2, you will go inside ChatGPT, Claude, and similar systems. How are they actually trained? What are tokens? What does it mean for a model to "hallucinate"? Why do these systems sometimes confidently say things that are completely false? And how do they scale to hundreds of billions of parameters?

Going Deeper

Want to go further on AI Ethics? Start here.

Reflect

Before you move on

No right answers here. These questions are for you.

A facial recognition system achieves 95% accuracy overall but only 73% accuracy on darker-skinned women. The company says the system is "highly accurate." What is wrong with that claim?

Aggregate accuracy hides unequal harm. A system that works well for the majority group and poorly for a minority can still report impressive headline numbers. The people most likely to be wrongly identified are already among the most vulnerable. Ethical evaluation requires disaggregating performance by demographic group, not just reporting a single score.

A hospital AI recommends denying a patient's insurance claim. The AI is a black box. Who is responsible if the decision turns out to be wrong?

Responsibility does not disappear because a machine made the recommendation. The hospital that deployed the system, the company that built it, and the clinician who accepted the output without scrutiny all bear some degree of accountability. AI does not create a responsibility vacuum; it redistributes it in ways institutions must think carefully about before deployment.

A highly accurate medical AI is also completely uninterpretable. A less accurate model can explain every decision in plain language. Which would you deploy in a hospital, and why?

There is no single right answer, and that is the point. A black-box model may save more lives on average, but its errors are invisible and impossible to contest. An interpretable model allows clinicians to catch mistakes, builds trust, and supports patient rights. Many jurisdictions require explainability for high-stakes decisions. The tradeoff is real, and the right choice depends on context, stakes, and who bears the consequences of error.

Progress
Done with this lesson?
Mark it complete to track your progress.